Repository navigation
fix(last9-cloudwatch): name the averaging window for MSK throughput - #18
Conversation
Confirming eval run 34454314035 caught a candidate reporting 200 instead of 250 bytes/second for a requested interval mean: it issued bare instant queries, saw only the latest period, and divided that pair. The grader failed it on value, evidence, window coverage, and raw observation. SKILL.md carries the weighted-window rule and the RDS, DMS, EC2, and ElastiCache references each name the latest-period versus whole-window distinction. The MSK throughput bullet said only "use the requested level or average", so it was the one family reference that left the choice open. Name both cases and say to read the raw companion series over the window rather than a single instant query. Refs ENG-1891. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The MSK failure was not only an MSK wording gap. SKILL.md gives the weighted-window formula in its statistics table, and gives the range selector separately, in a paragraph about inspecting raw timestamps. Nothing said that a selector carrying no range returns only the latest period, so a model could follow the formula correctly and still feed it a single period, which is what the failing run did. State it once in the tool reference, where query shape is already discussed, and give a self-check: the returned sample count must match the periods the interval should contain. This covers every family rather than the one whose reference text happened to be thinnest. Refs ENG-1891. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Verification result: the fix is unproven. Reporting it straight.I tested the hypothesis rather than assuming it. The runner accepts
The failure did not reproduce once in nine pre-fix attempts. Both arms are at ceiling, so this branch cannot be shown to change anything. My stated mechanism, a bare instant query with no range selector, did occur in the one CI failure, but it is not the model's normal behaviour even without the fix. So the honest position on the two commits:
I have corrected the "Verification" section of the description, which promised evidence this run did not deliver. Recommendation: merge as a documentation clarification with no efficacy claim, or close it if you would rather not carry unverifiable changes. I would merge, because both statements are independently correct and one removes a genuine inconsistency between family references, but I do not want the merge justified by a fix that the data does not support. Broader regression check (run 34465699752, twelve cases, |
Why
The confirming eval run after #17 merged (34454314035, pinned to
3b39308) had one candidate semantic failure:msk-broker-ingressreported 200 bytes/second where the task asked for the mean over the interval and the answer is 250.The candidate issued two bare instant queries with no range selector, received only the latest period, and divided that pair (400 / 2) instead of totalling the disjoint periods (2000 / 8). The grader failed it on
value,evidence,window_coverage, andraw_observation, which is the correct verdict.Not a regression from the merge: the eval manifest shows the candidate skill hashes, evaluator hashes, and fixture hash are byte-identical between
ca085f7and3b39308. Same inputs, different sampling. That case passed in 6 of 7 candidate executions.The gap
SKILL.mdalready carries the rule in its statistics table: "Window average = total Sum / total SampleCount. Averaging period averages is wrong when counts differ." Four family references restate the choice at the point of use:rds-aurora.mddms.mdec2.mdelasticache.mdmsk.mdMSK was the single family reference with a window-mean scenario that left the window unnamed. This aligns it with the other four rather than adding a new rule.
Change
One sentence in the throughput bullet of
references/msk.md: name both cases, and say to read the raw companion series over the window rather than a single instant query returning only the last period. Therate()warning and the broker/topic overlap warning are unchanged.scripts/check-skill-pack.shand the selftest pass.Verification
Nine candidate samples of the failing case against master's skill and nine against this branch: both 9/9 pass, both used a range selector in 9/9, both answered 250. The failure did not reproduce in the pre-fix arm, so this change is unproven and is offered as a documentation clarification, not a demonstrated fix. Details in the comments.
Refs ENG-1891.
🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.